Skip to main content
The following model families are supported. Download .gguf files from HuggingFace and pass them to --model.
DeepSeek-V3 and DeepSeek-R1 use Multi-head Latent Attention (MLA). Pass -mla 3 (the default) for best performance. Lower values reduce memory use at a speed cost.
Do not use Unsloth models with _XL in their name that contain f16 tensors. These models are incompatible with ik_llama.cpp. Unsloth _XL models that do not use f16 tensors are fine.